IEEE Transactions on Medical Imaging
● Institute of Electrical and Electronics Engineers (IEEE)
Preprints posted in the last 90 days, ranked by how well they match IEEE Transactions on Medical Imaging's content profile, based on 21 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Savant, V.; Wang, Y.; Xuan, J.
Show abstract
Voxel-level annotation for volumetric medical imaging is expensive and difficult to scale, which makes training highcapacity 3-D segmentation models challenging in practice. Transfer learning (TL) from large public datasets is a common remedy, but it can under-perform when the source domain differs from the target anatomy and acquisition characteristics, as is often the case for pulmonary nodules. In this work, we propose a masked autoencoder (MAE) pretraining-based approach to break the data efficiency wall of domain difference and present a focused empirical study of domain-specific self-supervised learning (SSL) for 3-D lung nodule segmentation. We evaluate two experimental settings: first, Masked Autoencoder (MAE) pretraining versus random initialization across representative baselines; second, MAE versus Decathlon TL for UNETR++ while testing whether MAE-based pretraining also benefits a CNN baseline (V-Net). MAE pretraining on target-domain CT volumes achieves a Dice Similarity Coefficient (DSC) of 0.307, outperforming random initialization (0.136) and Decathlon weights (0.257). In addition, MAE improves the stability of V-Net in a "low-data" regime (i.e., with "insufficiently labeled" data), increasing DSC from 0.010 to0.071. Overall, these results suggest that MAE-based pretraining can provide a practical and robust initialization strategy for volumetric segmentation when labeled data are limited.
Xie, C.; Hu, B.; Alakeel, A. M.; Fleischer, C. C.; Fedorov, A. G.
Show abstract
The development of digital twins in medicine, i.e., virtual replicas of human organs, offers a promising path toward precision medicine by enabling interpretable, mechanistic, and actionable insights. In the brain, cerebrovascular twins support individualized modeling of hemodynamics and bio-transport, with broad applications. A major bottleneck, however, is the lack of robust methods to transform in vivo cerebrovascular images into simulation-ready cerebrovascular meshes or graphs. Here, we present CerebroVascular Imaging to Graph reconstruction (CVIG), a robust and multiscale framework for reconstructing whole brain cerebrovascular graphs from in vivo cerebrovascular images. CVIG integrates vessel vectorization, with tolerance to discontinuity in vessel structures, using a topology-guided assembly of vessel trees to generate cerebrovascular graphs from medical images. We demonstrate the ability of CVIG to generate vascular graphs with improved vascular coverage and topological correctness, the capability essential for high fidelity brain biophysical simulations. This work establishes a vascular graph framework for individualized modeling and analysis, providing a key foundation for digital twins of the human brain.
Das, P.; Rath, J.; Jaiswal, A.; Dash, B. B.
Show abstract
Accurate brain tumor segmentation from magnetic resonance imaging (MRI) remains a challenging task because supervised deep learning models require large quantities of annotated data, which are expensive and time-consuming to obtain. This study investigates whether synthetic MRI images generated using a Deep Convolutional Generative Adversarial Network (DCGAN) can improve U-Net-based brain tumor segmentation using synthetic data augmentation. Experiments were performed on the LGG-MRI dataset comprising 3,929 image-mask pairs. A baseline U-Net was first trained using the original training dataset. Synthetic MRI images were subsequently generated using a DCGAN, and threshold-derived pseudo-masks were assigned to the generated images to construct an augmented training dataset. The same U-Net architecture was then retrained using the augmented dataset and evaluated on an identical held-out test set. Compared with the baseline model, DCGAN-based augmentation increased the Dice coefficient from 0.2067 to 0.3037 and the Intersection over Union (IoU) from 0.1243 to 0.1918, while reducing the final test loss from 0.0474 to 0.0275. These results indicate that synthetic MRI augmentation was associated with improved segmentation performance under the experimental conditions of this study. However, the reliance on threshold-derived pseudo-labels and evaluation on a single dataset limit the generalizability of the results. The proposed workflow provides a reproducible implementation for evaluating DCGAN-based synthetic data augmentation in supervised brain tumor segmentation and establishes a baseline for future studies employing more reliable annotation strategies and broader experimental validation.
Tecchio, P.; Schlaffke, L.; Bolsterlee, B.; Hahn, D.; Raiteri, B. J.
Show abstract
Muscle architecture shapes muscle function and changes with age, growth, training and disease, yet quantifying three-dimensional (3D) muscle architecture in vivo remains challenging. We introduce a hybrid fascicle tractography approach for freehand 3D ultrasound data that accurately reconstructs 3D muscle fascicles with respect to an objective, anatomically relevant coordinate system defined by the muscle's central aponeurosis. The hybrid approach combines Hessian-based fascicle detection with wavelet-based refinement to generate volumetric fascicle orientations. In a synthetic dataset with known ground truth, fascicle orientations and lengths were estimated with errors of [≤]2{degrees} and ~1.5%, respectively. In vivo, the approach detected physiologically plausible fascicle lengthening in the human tibialis anterior following a passive plantar flexion rotation, whereas diffusion tensor imaging of the same muscle did not. The proposed method enables anatomically relevant, objective and non-invasive quantification of 3D muscle architecture in vivo, providing a practical framework for applications in clinical and applied muscle physiology.
Cheng, M.; Liu, C.; Gu, L.
Show abstract
Class imbalance is a prevalent issue in medical image classification that significantly degrades a model's capacity to recognize minority-class lesions, thereby restricting its applicability in real-world clinical screening scenarios. Existing studies typically address this problem through data resampling, loss re-weighting, or decision boundary adjustment strategies; however, these methods predominantly focus on compensation during the classification stage. In contrast, the representation learning process in earlier stages is often dominated by easy majority-class samples, and its impact on the feature quality of minority classes has not received adequate attention. To address this issue, we propose an Imbalance-Aware Robust Representation Learning (IRRL) framework for class-imbalanced medical image classification. IRRL prioritizes the refinement of minority-class-related local representations before global classification. Specifically, implicit local token representations are constructed from convolutional feature maps based on their receptive-field structure. Semantic confidence-guided reliability estimation, difficulty-adaptive supervised contrastive learning, and minority-class prototype regularization are then introduced to improve the learning of informative local representations and hard minority-class samples. Finally, a Transformer performs global context modeling for image-level classification. Experiments on four public datasets, including ISIC 2018, PAD-UFES-20, OCTID, and BUSI, show that IRRL achieves balanced classification performance, with favorable F1-score and Matthews Correlation Coefficient (MCC) results that reflect improved minority-class recognition quality. The results across datasets with different imaging modalities and imbalance conditions further demonstrate the robustness and consistency of the proposed representation learning strategy.
Ilovitsh, T.; Shapiro, G.; Gershman, Y.; Bismuth, M.
Show abstract
This study presents the use of sub-micron nanobubbles (NBs) as contrast agents for ultrasound localization microscopy (ULM), a super-resolution imaging technique that visualizes microvascular structure and flow beyond the acoustic diffraction limit. While ULM has traditionally relied on micron-sized microbubbles (MBs), the reduced dimensions and prolonged circulation times of NBs make them attractive candidates for localization-based imaging. However, their weaker acoustic responses present significant challenges for reliable detection and tracking. To address this challenge, we developed the ULM Master GUI, an interactive framework for optimization of the complete ULM processing pipeline. Using custom ultrasound-compatible wall-less gelatin flow phantoms containing vessel-mimicking channels and bifurcations ranging from 100 to 500 m, we demonstrate that NB-based ULM achieves velocity reconstruction and flow partitioning measurements comparable to conventional MB-based ULM. Across all investigated geometries, NBs faithfully reproduced the underlying flow patterns and hemodynamic behavior despite their substantially reduced acoustic scattering. These findings establish the feasibility of NB-based ULM, expand the range of contrast agents available for localization microscopy, and provide a foundation for future super-resolution ultrasound imaging using nanoscale acoustic contrast agents. The ULM processing GUI is publicly available at https://github.com/grisha1998/ulm-super-resolution-toolbox.
Ye, Z.; He, F.; Zhao, T.; Xia, W.
Show abstract
Ultrathin endoscopy is highly attractive for real-time tissue imaging in narrow and hard-to-reach regions of the body. A single multimode fibre (MMF) is an attractive probe because of its small diameter, flexibility, and diffraction-limited spatial resolution enabled by the large number of transverse modes guided within a single core. Because the distal fibre tip is inaccessible during endoscopy, reflection-mode imaging, in which the same fibre delivers illumination and collects backscattered light, is more practical than transmission-mode imaging. However, image recovery from the resulting speckle pattern is challenging because light undergoes double-pass propagation through the MMF, with mode coupling and dispersion; the backscattered signal is weak, and the camera records intensity only, without phase information. Here, we propose a single-shot reflection-mode MMF imaging framework that combines a reflected real-valued intensity transmission matrix (reflected-RVITM) with an image restoration network. The reflected-RVITM is calibrated using intensity-only measurements, without interferometry or phase retrieval, and provides a physics-guided initial reconstruction from a single backscattered speckle frame. A restoration network then refines this initial reconstruction instead of inverting the raw speckle. Four restoration backbones are evaluated: HPM-Attention-UNet, GAM, MambaIRv2, and CICPNet. On matched datasets, hybrid models outperformed corresponding networks trained to map raw speckle directly to images. For example, HPM-Attention-UNet on MNIST improved mean PCC from 0.572 to 0.944 (+65.1%). Under domain shift, with training only on Fashion-MNIST and tested on unseen CIFAR scenes, hybrid models achieved mean PCC of 0.61-0.65, compared with 0.36-0.50 for direct learning. This framework is further demonstrated using physical objects at the distal fibre tip. These results demonstrate that a reflected-RVITM physics prior combined with a restoration network enables single-shot image recovery after intensity-only calibration, offering a phase-retrieval-free and generalisable route towards minimally invasive reflection-mode MMF endoscopy.
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.
Show abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Serial optical coherence tomography (S-OCT) is an imaging modality relying on the intrinsic contrast of a sample. When applied to brain tissues, the S-OCT contrast is primarily driven by the myelin reflectivity. Due to its high resolution, on the order of microns, and its 3D nature, S-OCT offers promise for studying WM connections at the microscale. However, while other microscopy imaging modalities have been shown to enable tractography, whether the reflectivity contrast from S-OCT supports the reconstruction of long-range WM fascicles at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We improve microscale orientation distribution functions (ODF) estimation by implementing a sliding-window formulation allowing the estimation of ODF at S-OCT resolution, and use apodized Dirac delta functions for reducing unwanted interference. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that using multiscale Frangi filters for estimating ODF outperforms structure tensor analysis. We also show that particle filtering tractography with anatomical constraints enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing the thalamocortical white-matter projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles visible at the microscale, and that these connections are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.
Show abstract
Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.
Hassan, M. W.; Crook, K.; Gi, Y. J.; Lee, J.; Hossain, M. M.
Show abstract
Objective: This study aims to develop and validate a quantitative, depth-resolved anisotropy imaging framework that extends ARFI-based focal degree-of-anisotropy (DoA) estimation into two-dimensional mapping by modeling the depth-dependent relationship between shear modulus ratio (SMR) and peak displacement ratio (PDR). Methods: We propose APRIL (Adaptive Polynomial Regression for anisotropy Imaging via ARFI-induced DispLacements), a framework for quantitative, depth-resolved DoA imaging that adaptively selects polynomial regression or shape-preserving spline interpolation based on excitation PSF asymmetry. Training data were generated using an LS-DYNA3D + Field II simulation pipeline in homogeneous transversely isotropic media (SMR 0.9-4.9). Testing included shifted SMRs under varied acoustic conditions and three heterogeneous inclusion configurations (anisotropic inclusion in isotropic background and vice versa). Experimental validation was performed in an in-vivo murine tumor model over the time, ex-vivo chicken breast, and tissue-mimicking gelatin phantoms, using a Verasonics system with an L11-5v transducer. Results: APRIL achieved depth-resolved SMR prediction errors below 9% over 10-30 mm, with highest accuracy in the focal region (MAE 2.3%, RMSE < 0.1) and stable performance across PSF transition zones. In heterogeneous phantoms, it reconstructed anisotropy maps with SSIM up to 86% and MPE below 7%, accurately delineating inclusion boundaries. Under acoustic parameter variations, mean absolute errors remained below 10%, demonstrating robustness to system and tissue heterogeneity. Conclusion: APRIL enables robust, two-dimensional anisotropy imaging beyond focal estimates. Significance: The method provides a physically grounded and generalizable framework for clinically viable anisotropy biomarkers in muscle, tendon, kidney, tumor and breast tissues.
Dong, S.; Guan, M.; Yang, L.; Liu, G.; Rominger, A.; Ren, W.; Ni, R.; Wei, X.
Show abstract
Clinical treatment planning of near-infrared (NIR) brain stimulation requires patient-specific light dosimetry to optimize fluence delivery to cortical targets. The gold-standard Monte Carlo (MC) photon transport forward solver is accurate but computationally expensive and non-differentiable for personalized inverse design across subjects. Here, we present a foundation-model (FM)-encoded, differentiable implicit-neural surrogate for the MC solver. A pretrained 3D MRI/CT foundation model, VISTA3D, is domain-adapted to head phantoms with known optical properties to encode the subject anatomy. Next, an implicit neural representation is used to predict light fluence at arbitrary continuous coordinates. This formulation enables off-grid queries and gradients with respect to illumination parameters. Trained with a physics-informed, decade-stratified loss, the surrogate attains R2 {approx} 0.90 on held-out subjects. Ablation results show that the FM benefit is contingent on domain adaptation. Benchmarked against standard learned surrogates, our model is the most accurate in the high-dose region and best on dose-fidelity metrics ({gamma}-index, treated-volume DICE). Finally, gradient-based optimization through the surrogate recovers MC-consistent illumination configurations 50-240 x faster.
Chattopadhyay, T.; Shelar, K.; Thomopoulos, S. I.; Thompson, P. M.
Show abstract
Scaling laws describe how model performance improves as the amount of training data increases, and recent theories such as the zeta law suggest that scaling behavior is influenced by the eigenspectrum of the models latent representation. Here, we evaluated whether the distribution of discriminative signals across spectral modes predicts the future scaling behavior, for MRI transformers trained for disease classification. We trained three supervised 3D vision transformers (ViT3D, MINiT, and NIT) for Alzheimers disease classification using 2,822 training scans from the Alzheimers Disease Neuroimaging Initiative (ADNI); we compared their encoder spectra with that of a frozen self-supervised DINO ViT-B/16 encoder adapted to 3D MRI. The supervised models learned highly concentrated representations, with 90-96% of CLS-token variance captured by a single principal component, whereas DINO distributed signal across many latent directions. Via spectral expansion of the Mahalanobis signal, we found that supervised training concentrated disease information into a single dominant mode, while self-supervised training produced a richer spectral geometry with higher effective rank and discoverability. This led to different scaling behavior: supervised models exhibited flatter AUC(N) curves, yet DINO continued to improve as sample size increased, gaining 11.0 percentage points from N=50 to N=2,822. Overall, the spectral distribution of the discriminative signal, for these different encoder types, influenced how much performance remained discoverable as sample size increased. Distributed representations may retain signal across many latent modes and continue to improve with additional data, whereas concentrated representations tend to exhaust most of the discoverable signal at much lower sample sizes.
Trisha, S. M.; Rahman, M. A.; Hassan, M. W.; Gi, Y. J.; Lee, J.; Hossain, M. M.
Show abstract
Viscoelastic characterization of tissue has significant diagnostic value in oncology, as tumor progression alters both elasticity and viscosity in ways that neither property alone can fully capture. Existing acoustic radiation force (ARF)-based methods such as Viscoelastic Response (VisR) ultrasound estimate relative elasticity and viscosity through per-A-line nonlinear model fitting, which is computationally intensive and requires auxiliary simulations to correct elasticity-dependent bias. This work presents VESTA (Machine Learning-Enabled Estimation of ViscoElastic Ratios from On-Axis Spatio-Temporal ARFI Features), a two-stage data-driven pipeline that predicts elasticity ratio (ER) and viscosity ratio (VR) directly from seven normalized ARFI displacement features at the A-line level, without model fitting or compensation. Stage~1 is an MLP classifier that detects inclusion boundaries from normalized peak displacement and negative peak velocity ratios; Stage~2 is a dilated Conv1D regression model that estimates ER and VR along the full axial sequence using the predicted mask alongside displacement features. The pipeline was trained on 500 simulated inclusion scenarios spanning three geometries, five focal depths, two F-numbers, and a broad range of material contrasts. In silico, mean predicted ER and VR were within 12\% of ground truth across all geometries, with performance best when ER and VR were moderate or decoupled. Experimental validation on a chicken breast phantom demonstrated plausible generalization to real tissue heterogeneity. Applied to an in vivo murine 4T1 breast cancer model, the pipeline tracked treatment-related attenuation of mechanical contrast in paclitaxel-treated tumors relative to controls over a 36-day imaging period, supporting its relevance for tumor monitoring.
Chen, W.-Y.; Wan, S.-Y.; Lin, G.-Y.
Show abstract
Accurate segmentation of thin-wall organs-at-risk (OARs)-the cochlea, vestibular semicircular canals, internal auditory canal, tympanic cavity, and middle ear-is clinically relevant for head-and-neck radiotherapy planning, yet these small, thin-wall structures remain among the most challenging targets for automated delineation. Dual-frequency feature fusion is a promising direction for boundary-sensitive representation, but under the investigated FP16 FFT-FcaNet setting, we observe an approximately 863-fold activation-scale mismatch between the FFT and FcaNet branches, causing a nominal 5 percent residual coefficient to behave as an approximately 43-fold dominant term. We propose FreqFuseNet, which resolves this mismatch by normalizing the FcaNet branch to the FFT activation scale before residual injection with a fixed low-amplitude coefficient (beta = 0.05), restoring beta as an interpretable 5 percent residual-amplitude coefficient relative to the FFT feature scale. Under a controlled binary per-OAR ROI protocol on the SegRap2023 head-and-neck CT benchmark across 10 clinically prioritized thin-wall OARs, FreqFuseNet achieves Dice of 0.849, HD95 of 0.824 mm, and SDice@1mm of 0.959 in the primary seed, with comparable performance in an independent second seed (Dice 0.843, HD95 0.823 mm). FreqFuseNet yields statistically significant case-level aggregate improvements over 3D U-Net and MedNeXt-S (Wilcoxon p < 0.01 and p < 0.05, respectively), using only 29.7 million parameters versus 414.6 million for the full wavelet baseline.
Partridge, T.; Ahmad, R.; Astolfo, A.; Buchanan, I.; Endrizzi, M.; Hawkins, M.; Olivo, A.; Esposito, M.
Show abstract
Quantifying cells within intact three-dimensional biological specimens remains a major challenge, as standard optical and histological techniques are inherently two-dimensional, destructive, or constrained by light scattering. Optical clearing can extend imaging depth but is time-consuming, disruptive to tissue integrity, and often incompatible with downstream analyses, limiting its practical use for routine three-dimensional quantification. X-ray computed tomography can overcome these limitations, yet conventional micro-CT lacks the soft-tissue contrast required for cellular-scale analysis. Here, we introduce an integrated imaging framework in which propagation-based phase-contrast X-ray CT is combined with volumetric nuclear segmentation to enable three-dimensional cell quantification in unstained volumetric tissue. We imaged ex vivo human liver tissue and segmented nuclei throughout the reconstructed volume, extracting quantitative nuclear metrics and spatial organisation metrics, including equivalent diameter, minor-to-major axis ratio and nearest-neighbour distance. We assessed measurement consistency across two non-overlapping volumes of interest and benchmark slice-resolved nuclear metrics against haematoxylin and eosin histology. The resulting high-contrast volumetric datasets preserve tissue context, allowing quantitative measurements to be interpreted alongside surrounding architecture and microstructure. Together, these results show that laboratory phase-contrast X-ray CT supports nucleibased volumetric cell quantification in intact unstained tissue and provides a framework for context-preserving quantitative analysis in three dimensions.
Pan, Y.; Feng, Y.; He, J.; Consagra, W.; Westin, C.-F.; Rathi, Y.; Ning, L.
Show abstract
Diffusion MRI (dMRI) enables noninvasive characterization of white-matter fiber orientations and tissue microstructure, but widely used approaches, such as constrained spherical deconvolution (CSD) and parametric multicompartment models, typically address these features separately. The diffusion tensor distribution (DTD) framework jointly represents fiber orientation and microstructure, but estimating DTD from finite, noisy measurements is severely ill-posed. Existing inversion methods either rely on nonnegativity constrained basis representations, which are challenging to sale to high-dimensional and high-resolution distributions, or use sampling-based approaches with limited reliability. We propose MaxEnt-DTD, a maximum-entropy algorithm for DTD estimation from finite and noisy dMRI data. By deriving the Lagrange dual formulation, we reformulate a constrained infinite-dimensional optimization problem into a finite-dimensional unconstrained convex optimization problem, substantially reducing the parameter space and enabling tractable whole-brain DTD estimation. We evaluate MaxEnt-DTD using both synthetic and in vivo data from the Human Connectome Project protocol and a second dataset using advanced B-tensor diffusion encoding. We compare MaxEnt-DTD-derived fiber orientation distributions with results from CSD and Monte-Carlo inversion methods, and assess fiber-specific microstructure measures and rotation-invariant metrics based on the cumulants of DTD. The results demonstrate that MaxEnt-DTD provides a reliable and efficient framework for joint fiber-orientation and microstructure analysis in dMRI.
Shenoy, A. R.; Mendez, T.
Show abstract
Stroke is a leading cause of death and long-term disability worldwide, affecting approximately 15 million individuals annually. Prompt and accurate subtype differentiation between ischemic and hemorrhagic stroke is clinically critical, as the two conditions demand diametrically opposite interventions - thrombolytic therapy versus surgical decompression. Yet the majority of existing deep learning approaches reduce this problem to binary detection, and virtually none address the opacity of their decision-making in a clinically actionable manner. We present CerebAI, an explainable, deployment-oriented three-class CT stroke classification system built on a fine-tuned ConvNeXt-Base backbone with Integrated Gradients (IG) attribution. Trained on 6,774 non-contrast CT scans stratified across No Stroke, Ischemic Stroke, and Hemorrhagic Stroke, CerebAI achieves a weighted F1-score of 0.9746 (95% CI: [0.9625, 0.9851]), accuracy of 97.47%, macro-averaged AUC of 0.9921, mean Intersection-over-Union (mIoU) of 0.9276, Expected Calibration Error (ECE) of 0.0115, mean Brier Score of 0.0150, and Cohen's {kappa} of 0.9483 - surpassing ResNet-50, EfficientNet-B4, and Vision Transformer (ViT-B/16) baselines across all reported metrics. Integrated Gradients produce pixel-precise saliency maps that localize pathological regions with greater anatomical fidelity than Gradient-weighted Class Activation Mapping (Grad-CAM), a finding we support with side-by-side qualitative comparison. CerebAI additionally incorporates a native DICOM processing pipeline to facilitate future clinical translation. Code and model weights are publicly available to support reproducibility and further research.
Kesenci, Y.; Le Folgoc, L.; Angelini, E.
Show abstract
Deep-learning-based segmentation algorithms have gained considerable accuracy for processing biological images. In particular, the introduction of large foundation models, novel architectures, and semantically varied datasets now allows for deployment of state-of-the-art models for clean image cohorts with limited re-training or, in the best of cases, in an out-of-the-box fashion. Biological imaging, however, is liable to corruptions that can hinder their deployment. While some methods document their robustness to the most common corruptions, a systematic robustness analysis of the state of the art to the expansive gamut of corruptions in biological imaging remains to be done. We perform this benchmarking by simulating 36 corruption types with varying degradation severity on images sampled from 30 different datasets. Our benchmark accounts both for the variety in biological images and the nature of corruptions. Among other things, our study reveals that performance on clean images does not correlate with overall robustness to image corruptions. In fact, we find that a decade-old method, StarDist, is more robust than many of its more recent foundation-model-based counterparts. We also show in a dedicated representation analysis that the performance of segmentation models collapses in the early layers of the encoding phase.
Sigger, N.; Nguyen, T. T.; Ashraf, S.; Tozzi, G.
Show abstract
Hyperspectral imaging (HSI) has gained increasing attention for bone assessment because it captures rich wavelength dependent information associated with mineralised tissue. HSI provides detailed spectral information related to material composition, while 3D geometric information supports the analysis of surface morphology and structural detail. However, integrating spectral and geometric information remains challenging, particularly when conventional reconstruction pipelines depend on external pose estimation. To address this challenge, we propose BoNeRF-HS, a self-calibrated hyperspectral neural radiance field for 3D reconstruction. BoNeRF-HS jointly optimises camera intrinsics, volume density, and hyperspectral radiance, removing the need for COLMAP based poses. To improve spectral modelling, we incorporate a gated spectral adapter head that learns wavelength dependent radiance features for hyperspectral view synthesis. We evaluate BoNeRF-HS on a multi-view hyperspectral dataset containing mouse bone, trabecular bone analogue, and cortical bone analogue samples. Experimental results demonstrate that our framework achieves improved reconstruction quality, and better preservation of bone surface details compared with existing approaches.
Wang, N.; Abraham, D.; Shah, Z.; Lin, Y.; Cao, X.; Wu, H.; Polimeni, J.; Huber, R.; Liu, Q.; Ning, L.; Rathi, Y.; Westin, C.-F.; Mattern, H.; Speck, O.; Yang, B.; Abad, N.; Liao, C.; Kerr, A.; Setsompop, K.
Show abstract
Purpose: To develop a Field-Correcting GRAPPA (FCG) technique to correct the spatiotemporal-varying phase errors in EPI caused by eddy currents. Methods: The fast-changing gradient in EPI causes strong eddy current effects and associated spatiotemporal-varying phase errors, producing significant image artifacts. The use of higher gradient amplitude, slew rate, and ramp sampling factor for faster imaging exacerbates this problem. In this work, FCG was developed to address this challenge by using a multi-layer perceptron (MLP) to provide a compact representation of a family of GRAPPA-like kernels that correct the spatiotemporal-varying phase errors in the data. A dedicated calibration pipeline was designed to acquire high-quality source and target data for MLP training in both slice-by-slice and simultaneous multi-slice (SMS) acquisitions. To validate FCG's assumptions and performance, a field camera was used to provide ground-truth measurement of phase patterns. The performance of FCG was further validated on phantom and in vivo experiments using demanding EPI trajectories across multiple 3T and 7T systems. Results: Field camera measurements revealed strong spatiotemporal phase variations along the kx direction that repeat along ky during EPI readouts. The experiments on high-performance systems across 3T and 7T demonstrate that FCG can provide superior correction for the artifacts induced by spatiotemporal-varying phase errors compared with existing approaches. Conclusion: FCG is an effective and robust method for correcting spatiotemporal phase errors in EPI, enabling improved image quality on high-performance systems.